# Overview Hacker’s Guide: [https://acmucsd.notion.site/diamondhacks-hacker-guide-2026](https://acmucsd.notion.site/diamondhacks-hacker-guide-2026) Browser Use requirements: ### **πŸ“‹ Requirements** * **Meaningful Impact:** The project must deliver on its promise to its target audience and provide a compelling experience. * **Browser Use Integration:** The core functionality must rely on Browser Use agents actively interacting with web environments. * **Working Prototype:** The project must be fully software based and functional for a live demo. Fetchai requirements: ### **πŸ“‹ Requirements** * **Agent Orchestration:** Develop a single or multi-agent orchestration that demonstrates reasoning, tool execution, and solves a real-world problem. * **Framework Flexibility:** Use any agentic framework like the Claude SDK, OpenAI Agent SDK, Google ADK, LangGraph, CrewAI, or simple plain Python to bring your idea to life. * **Agentverse Registration:** Register your agents with [Agentverse](https://agentverse.ai) and implement the Chat Protocol (mandatory) and Payment Protocol (optional) to support direct ASI:One interactions and built-in monetization. * **ASI:One Demonstration:** No custom frontend is required. The use case must be demonstrated directly through [ASI:One](https://fetch.ai/?utm_source=diamondhacks2026&utm_medium=direct&utm_campaign=hackathon). ### **πŸ“€ Deliverables** * **ASI:One Chat Session:** Share your ASI:One Chat session URL. * *Example:* `https://asi1.ai/shared-chat/6d7caa05-94ff-4d33-8fa1-8e24d8ec032e` * **Agentverse Profile:** Provide the URL for your Agent(s) on Agentverse. * *Example:* `https://agentverse.ai/agents/details/agent1qtuuyttz8ujuxceq0gllcerlksjneenrh2mfcm67st8qrm9lzzh3cd7f9h6/profile` * **Source Code & Demo:** A public GitHub repository link and a demo video submitted on Devpost. TwelveLabs Judging Criteria: ### **πŸ“‹ Judging Criteria** * **Twelve Labs API Use**: Core functionality driven by our video-understanding endpoints * **Impact & Usefulness**: Real-world relevance for media entertainment, public sector, and marketing advertising workflows * **Wow Factor**: Creativity and uniqueness of your solution * **Technical Depth**: Complexity and polish of your implementation * **UX Quality**: Ease of use & presentation JellyJelly Judging Criteria: **πŸ“‹ Judging Criteria** * **JellyJelly API Integration:** How effectively is our API used to power core features (clipping, AI-summarization, or wallet interactions)? * **Economic Innovation:** Creative use of the $500 stablecoin reward structure or JELLYJELLY token utility within the app. * **User Experience (UX):** Is the tool intuitive for creators? Does it make the "capture-to-post" flow frictionless? * **Technical Depth:** Complexity of the implementation, especially in handling video data or Solana-based transactions. * **The "Vibe" Factor:** Does it align with the fun, community-driven spirit of the JellyJelly ecosystem? ## **Judging πŸ§‘β€βš–οΈ** Projects will be judged in four main criteria: **Idea, Experience, Implementation,** and **Demo/Presentation**. After the project submission deadline, participants will showcase their projects in a science-fair style format, where judges will visit each presentation. More details will be announced on Discord closer to the deadline. # Tracks # **DiamondHacks 2026 β€” Complete Track Guide** Teams can win multiple awards across categories. You can win an Overall Prize, a Sponsor Prize, AND a Track Prize β€” but you're limited to one main track since you can only submit to one. --- ## **Main Tracks (Choose Up to One)** ### **1\. Alchemy of the Earth (Social Impact / Sustainability)** **Prize:** Owala 40 oz Bundle Build solutions for the world's most pressing environmental and social issues. Projects should focus on sustainability, climate change, or empowering underserved communities. Think conservation tech, carbon footprint tools, environmental monitoring, resource optimization, community empowerment platforms, or green energy solutions. The key theme is using technology as "green magic" to protect the planet and ensure a sustainable future for all. --- ### **2\. Elixirs of Vitality (Health)** **Prize:** Peak Design Tech Pouch Dedicated to healing and well-being. Projects should improve physical health, mental clarity, or medical accessibility. This covers a wide range β€” vitals monitoring, mental health platforms, telemedicine tools, fitness tracking, medical data accessibility, wellness apps, or assistive health technology. The emphasis is on preserving life and making healthcare more accessible. --- ### **3\. The Scholar's Spellbook (Education)** **Prize:** Mini Projector Reimagine how knowledge is shared and learned. Build immersive learning platforms, interactive educational tools, or gamified tutorials that make complex subjects intuitive. Think AI tutors, interactive textbooks, language learning tools, STEM visualizers, study aids, or platforms that democratize education. The goal is making education engaging and accessible for everyone. --- ### **4\. Enchanted Commerce (Commerce / E-Commerce)** **Prize:** Apple AirTags Revolutionize how we exchange goods and manage money. Build seamless payment systems, immersive virtual storefronts, logistics optimization tools, or clever marketplace innovations. This includes anything from personalized shopping experiences to supply chain management, price comparison engines, or novel checkout flows. Turn shopping into a standout experience. --- ### **5\. The Rogue's Ritual (Wildcard)** **Prize:** Sony SRS-XB100 Speaker For projects that defy categorization. If your idea doesn't fit neatly into sustainability, health, education, or commerce, this is the track. Revolutionary tools, creative experiments, boundary-breaking concepts β€” anything goes. The most powerful projects here are the ones that create an entirely new category. --- ## **Sponsor Tracks** ### **6\. Best Interactive AI β€” *Sponsored by UCSD CSE*** **Prize:** M4 Apple iPad Air 13-inch **What is this:** This is UCSD's own Computer Science department sponsoring a track. There's no specific product or API to use β€” they just want to see AI applied to entertainment in an interactive way. **What you'd be doing:** Taking an existing app, game, media tool, or text interface and bolting on an AI feature that dynamically responds to user input. The emphasis is on *fun and engagement*, not utility. They explicitly encourage using starter templates or open-source codebases so you don't waste time on UI β€” just focus on the AI integration itself. Use AI to build an interactive software project focused on **entertainment**. The goal is to take a standard application, text interface, or media tool and use AI to make it dynamically respond to user input, making it more engaging and fun. **Requirements:** * Interactive entertainment focus β€” something people can actively play with or experience for fun * Integrate at least one "smart" AI-driven feature as the main interaction method * Working software prototype for a live demo * Encouraged to use starter templates or open-source codebases (focus on your AI integration, not building UI from scratch) **Deliverables:** * Two-sentence feature description of your AI feature, how it works, and why it's engaging * Working demo or video walkthrough * Public GitHub repo --- ### **7\. Best Use of Browser Use β€” *Sponsored by Browser Use*** **Prize:** 2 iPhone 17 Pros, 2 AirPods Max, and a week-long trip to the Browser Use Hacker House in San Francisco (for the winning team) **Merch:** Show something you built with Browser Use at their table to get merch. **What is Browser Use:** An open-source Python library (81k+ GitHub stars) that lets AI agents *control a real web browser*. It uses Playwright under the hood. Your AI agent can see a webpage (both visually and via HTML structure), identify interactive elements, click buttons, fill forms, navigate between pages, manage multiple tabs, and execute complex multi-step web workflows. It's model-agnostic (works with OpenAI, Anthropic, etc.). Think of it as giving your LLM actual hands to use the internet like a human would. **What you'd be doing:** Building AI agents that autonomously interact with real websites to accomplish tasks. Not just scraping β€” actually clicking, typing, navigating, and completing end-to-end workflows. Examples: agents that fill out job applications, book travel, submit forms, compare products across stores, or automate any multi-step web task. Your agent should *do the thing*, not just recommend it. Build an impactful project using Browser Use. "Impact" doesn't strictly mean utility β€” the main judging question is "How well does the product deliver for its intended audience?" Agents that shop, fill out applications, submit quizzes, automate web workflows β€” the web is your playground. **Requirements:** * Meaningful impact: must deliver on its promise to its target audience * Core functionality must rely on Browser Use agents actively interacting with web environments * Working software prototype for a live demo --- ### **8\. Best Use of [Fetch.ai](http://Fetch.ai) β€” *Sponsored by Fetch.ai*** **Prizes:** * 1st Place: $300 cash (split among team) * 2nd Place: $200 cash (split among team) **What is Fetch.ai / Agentverse / ASI:One:** Fetch.ai is building infrastructure for the "agentic web." There are three pieces: * **Agentverse** is an open directory/platform where you deploy, host, and register AI agents. It has 2M+ agents. Think of it as an app store for AI agents. * **ASI:One** is the consumer-facing AI that orchestrates those agents. A user makes a request ("plan a trip to NYC with flights, hotels, and dinner") and ASI:One discovers the right agents on Agentverse and coordinates them to complete the task. * **Chat Protocol** (mandatory) lets your agent communicate through ASI:One. **Payment Protocol** (optional) lets agents transact. **What you'd be doing:** Building one or more AI agents using any framework (Claude SDK, OpenAI, LangGraph, CrewAI, plain Python, etc.), registering them on Agentverse, and demonstrating them through ASI:One. Your agents should turn a user's natural language intent into real, executed outcomes β€” not just information retrieval. The key is: a user talks to ASI:One, your agents get discovered, and they *do something real*. Build and register AI Agents on Agentverse, discoverable via ASI:One, that turn user intent into real, executable outcomes. **Requirements:** * Agent orchestration: single or multi-agent system demonstrating reasoning, tool execution, and solving a real-world problem * Framework flexibility: use any agentic framework (Claude SDK, OpenAI Agent SDK, Google ADK, LangGraph, CrewAI, plain Python, etc.) * Agentverse registration: register agents with Agentverse, implement the Chat Protocol (mandatory) and Payment Protocol (optional) * ASI:One demonstration: no custom frontend required β€” demonstrate through ASI:One directly **Deliverables:** * ASI:One Chat session URL * Agentverse profile URL for your agent(s) * Public GitHub repo \+ demo video on Devpost --- ### **9\. Qualcomm Multiverse β€” *Sponsored by Qualcomm*** **Prize:** Meta Quest 3 512 GB for each member of the winning team Build intelligent experiences spanning AI PCs, mobile devices, cloud services, and microcontrollers. Distribute perception, intelligence, and interaction across platforms powered by Snapdragon. Use an AI PC as a central control surface, pair with mobile for real-time sensing, extend with cloud AI, and connect the physical world using Arduino UNO Q and/or Rubik Pi. **Requirements:** * Cross-device AI spanning PC, mobile, cloud, and microcontrollers * Intelligent workflows distributing inference and control across devices * Hardware integration with AI PC, mobile, Arduino UNO Q / Rubik Pi * Working prototype that's privacy-first, low-latency, and energy-efficient **Note:** Requires a separate application via Microsoft Form. --- ### **10\. Best Use of TwelveLabs β€” *Sponsored by TwelveLabs*** **Prizes:** * 1st Place: 50 hours of video indexing credits ($270 value) * 2nd Place: 40 hours of video indexing credits ($217 value) * 3rd Place: 30 hours of video indexing credits ($164 value) **What is TwelveLabs:** A video understanding API platform with two core models: * **Marengo** is an embedding model β€” it watches video (visuals, audio, text) and produces vector embeddings for search. You can search across massive video libraries with natural language queries like "find the scene where the car crashes into the wall" and it pinpoints the exact timestamp. It's "any-to-any" search: text-to-video, video-to-video, etc. * **Pegasus** is a generative model β€” it watches video and produces human-readable text. Summaries, highlights, answers to questions about the video content, chapter breakdowns, etc. Supports videos up to 1 hour long. Together, these let you build: semantic video search engines, auto-generated summaries/captions, video Q\&A systems, content moderation tools, video analytics dashboards, and any app that needs to "understand" what's happening in a video. **What you'd be doing:** Using the TwelveLabs API (Python/JS SDK) to index videos with Marengo and/or generate text outputs with Pegasus. Your project's core functionality should be powered by TwelveLabs' video understanding β€” not just using it as a minor feature. The API handles the heavy ML lifting; you focus on the application layer and UX. Harness state-of-the-art video foundation models (Pegasus and Marengo) to make sense of video. Search, summarize, analyze, or build entirely new video experiences. **What You Can Build:** * Semantic video search: find exact moments using natural language * Automated video summarization and captions for long-form or short-form content * Custom video analytics: extract insights, Q\&A bots, or context-aware tools across hours of footage **Judging Criteria:** * TwelveLabs API usage (core functionality driven by their endpoints) * Impact and usefulness (real-world relevance) * Wow factor (creativity and uniqueness) * Technical depth (complexity and polish) * UX quality (ease of use and presentation) **Resources:** * Developer Hub: [https://www.twelvelabs.io/developer-hub](https://www.twelvelabs.io/developer-hub) * API Docs: [https://docs.twelvelabs.io/docs/get-started/introduction](https://docs.twelvelabs.io/docs/get-started/introduction) * Sample Apps: [https://www.twelvelabs.io/sample-apps](https://www.twelvelabs.io/sample-apps) * Blog: [https://www.twelvelabs.io/blog](https://www.twelvelabs.io/blog) --- ### **11\. Best Use of Jelly API β€” *Sponsored by JellyJelly*** **Prizes:** * $500 in stablecoins (split among winning team) * Summer fellowship opportunity for every member of the winning team **What is JellyJelly:** A social network app (founded by Venmo co-founder Iqram Magdon-Ismail) built on Solana. The core concept: users have video chats and livestreams, and the app lets you capture, caption, and share clips from those calls β€” called "Jellies." It has built-in AI for auto-captioning and titling. The JELLYJELLY token (a Solana SPL token) powers the economy: tipping creators, boosting content visibility, unlocking premium features, and waitlist skipping. Think of it as "TikTok meets Venmo on Solana" β€” raw, candid video clips with decentralized monetization baked in. **What you'd be doing:** Using the JellyJelly API to build tools around their video clip \+ monetization ecosystem. This could mean: AI-powered auto-clipping of livestreams, new ways to monetize content with JELLYJELLY/USDC, bots or overlays for live video, or cross-platform tools that export Jellies to other social networks. You're building on top of their existing social/video/crypto infrastructure. JellyJelly is a social network built on Solana that transforms video chats and live streams into shareable, monetizable "Jellies." Build the future of SocialFi and AI-driven content. **What You Can Build:** * AI video clipping and curation from long-form video chats or streams * SocialFi monetization tools (paywalled content, automated tipping, revenue sharing with JELLYJELLY or USDC) * Interactive "Jelly" layers (bots/overlays for real-time AI polls, summaries, RSVP events from video metadata) * Cross-platform social bridges (export/optimize Jellies for TikTok, Instagram, X with AI-generated captions) **Judging Criteria:** * JellyJelly API integration effectiveness * Economic innovation (creative use of stablecoin rewards or JELLYJELLY token utility) * User experience (intuitive for creators, frictionless capture-to-post flow) * Technical depth (video data handling, Solana transactions) * "Vibe" factor (alignment with JellyJelly's community-driven spirit) --- ### **12\. MLH Sponsored Tracks** Details available at: [https://www.mlh.com/events/diamondhacks-2026-a6/prizes](https://www.mlh.com/events/diamondhacks-2026-a6/prizes) Includes tracks like Best Use of Gemini API, Best Use of ElevenLabs, Best Use of Solana, Best Use of Vultr, and Best .Tech Domain Name. **Notable MLH Prizes:** * **Best Use of Gemini API:** Google Swag Kits * **What is Gemini:** Google's flagship multimodal AI model family (currently Gemini 3 series). The API gives you: text generation, multi-modal understanding (text \+ image \+ video \+ audio input), function calling, grounding with Google Search and Google Maps, structured JSON outputs, file/data processing, image generation (via Imagen 4), text-to-speech, and a Deep Research agent that autonomously plans and executes multi-step research. It's model-agnostic in the sense that you call it via REST or their SDK (Python/JS). * **What you'd be doing:** Using the Gemini API as the reasoning/generation backbone of your app. Could be a chatbot, a multi-modal analyzer, a research tool, a content generator, or anything that needs strong language/vision understanding. Free tier available. * **Best Use of ElevenLabs:** Wireless Earbuds * **What is ElevenLabs:** The leading AI voice platform. Their API does: text-to-speech (32 languages, emotionally expressive, natural-sounding), instant voice cloning (from 60 seconds of audio), professional voice cloning (from 30+ minutes), a library of 10k+ pre-made voices, speech-to-text transcription, voice changing, voice isolation, AI sound effects generation, music generation, and dubbing/translation. * **What you'd be doing:** Adding voice to your project via their API. Text-to-speech for narration, voice cloning for personalization, sound effects for immersion, or speech-to-text for input. The SDK is simple β€” a few lines of code to generate speech from text. Hackathon codes available at their table. * **Best Use of Solana:** Ledger Nano S Plus * **What is Solana:** A high-performance Layer 1 blockchain. Thousands of transactions per second, sub-second finality, near-zero fees (\~$0.00025 per transaction). Smart contracts are written in Rust (or via the Anchor framework for easier development). The ecosystem includes DeFi (DEXes, lending), NFTs, gaming, payments, DAOs, and real-world asset tokenization. * **What you'd be doing:** Building any on-chain component with Solana β€” token creation, payments, NFT minting, smart contract logic, decentralized identity, or integrating with existing Solana DeFi protocols. You'd use Solana's devnet for testing and demo. Anchor \+ Rust for smart contracts, or web3.js/Solana SDK for frontend integration. * **Best Use of Vultr:** Portable Screens * **What is Vultr:** A cloud infrastructure provider (like a simpler AWS). They offer: on-demand VPS servers, bare metal servers, Cloud GPUs (NVIDIA H100, GH200, A-series; AMD MI355X), managed Kubernetes, S3-compatible object storage, serverless inference, and global deployment across 32 data centers. Pay-as-you-go pricing. * **What you'd be doing:** Hosting your project's backend, running GPU-accelerated ML inference, deploying containers, or using their serverless inference for AI model serving. They're giving free cloud credits for the hackathon β€” claim them and use Vultr as your infrastructure regardless of whether you're targeting this prize specifically. * **Best .Tech Domain Name:** Desktop Microphone \+ free .Tech domain for up to 10 years * **What this is:** Register a .tech domain name for your project. That's it. Basically a free prize β€” just pick a cool domain name. Do this regardless of what you build. --- ## **Side Tracks (Stackable β€” Win These on Top of Other Prizes)** ### **13\. Best Solo Hack** **Prize:** Apple Watch SE 3 Build something incredible entirely on your own. --- ### **14\. Best Duo Hack** **Prize:** 27-inch Monitor Showcase the power of pair programming and tight collaboration as a two-person team. --- ### **15\. Best Beginner Hack** **Prizes:** * 1st Place: Razer Basilisk Hyperspeed * 2nd Place: Monitor Light Bar * 3rd Place: Anker Power Bank Perfect for first-time hackers. --- ### **16\. Best AI/ML Hack** **Prize:** JBL Go 3 Push the boundaries of artificial intelligence and machine learning. --- ### **17\. Best UI/UX Hack** **Prize:** Philips Hue 2-Pack For the project with the most beautiful and intuitive user experience. --- ### **18\. Best Mobile Hack** **Prize:** Massage Gun Bring your ideas to the palm of your hand with a standout mobile app. # Brainstorm FETCH MUST MONETIZE AGENT USING STRIPE \- flipping works well Google \- batch request optimization .tech domain name \- use diamondhacks code for free domain Lovable pro code: COMM-ACMU-2390 * an AI-powered browser copilot that organizes your tabs and workflows around what you’re trying to do. Instead of manually juggling tabs and repeating tasks, you can say β€œI’m studying physics” or β€œapply to internships,” and it automatically restores the right workspace and executes relevant browser actions. It learns and reuses workflows (β€œskills”) so your browser adapts to you over time. * An autonomous multi-agent system that turns photos of items into profit by identifying products, verifying real resale demand across marketplaces through live browser interaction, and automatically creating listings when a profitable opportunity is detected. Instead of guessing value, it executes the entire flipping workflow end-to-endβ€”from analysis to listingβ€”using coordinated agents. * Drop in any lecture video. Take notes as you watch. We semantically indexes the video with TwelveLabs, compares your notes against what was actually covered, detects the gaps, and has your professor explain exactly what you missed β€” in their own cloned voice. * Built on TwelveLabs (semantic video understanding), Gemini (gap reasoning \+ research), and ElevenLabs (professor voice clone). No more rewatching entire lectures. Just the parts you actually need. * RL agent optimizes city traffic light timing to minimize emissions β€” visual sim, watch it converge live * An AI system that learns to control traffic lights in real time to reduce congestion and emissions. Using a real city map simulation, it continuously adjusts signal timing and visibly improves traffic flow while lowering COβ‚‚ output. Instead of fixed schedules, it learns and adaptsβ€”showing live how smarter control can make cities cleaner and more efficient. # Flipping PRD # **PRD β€” \[FILLER\] Autonomous Resale Agent** **DiamondHacks 2026 | April 5–6 | UCSD** --- ## **1\. Overview** An autonomous multi-agent system that turns a photo of any thrift store item into a complete resale listing. The user uploads a photo. Four coordinated AI agents identify the item, research real market comps, compute profit margin, generate a clean product photo, and populate a ready-to-post Depop listing β€” fully automatically. The user clicks one button to post. **One-liner:** *Photo in. Listing out. Profit shown.* --- ## **2\. Problem** Thrift store flipping is a proven way to make money, but the research and listing process is tedious, time-consuming, and requires expertise most people don't have. People leave money on the table because they don't know what items are worth, what sold comps look like, or how to write a compelling listing. This system eliminates every manual step between finding an item and posting it for sale. --- ## **3\. Target User** A thrift store shopper standing in an aisle, phone in hand, who wants to know in 60 seconds whether an item is worth buying to flip β€” and if so, have the listing created automatically. --- ## **4\. Tracks** | Track | Type | Justification | | ----- | ----- | ----- | | Enchanted Commerce | Main | End-to-end resale automation, novel marketplace UX | | Best Use of Browser Use | Sponsor | Load-bearing: eBay comp scraping \+ Depop form population | | Best Use of Fetch.ai | Sponsor | 4 uAgents registered on Agentverse, orchestrated via Chat Protocol | | Best Use of Gemini | Sponsor (MLH) | Gemini Vision for item identification | | Best AI/ML Hack | Side | Multi-agent AI pipeline, vision model, semantic pricing | | Best UI/UX Hack | Side | Polished real-time agent activity dashboard | | Best .Tech Domain | Side | Register domain, free prize | --- ## **5\. Tech Stack** | Layer | Technology | | ----- | ----- | | Frontend | Next.js 15, shadcn/ui, Tailwind CSS | | Real-time updates | SSE (Server-Sent Events) β€” backend to frontend | | Backend | Python (FastAPI) | | Agent framework | Fetch.ai uAgents | | Agent registry | Agentverse (Chat Protocol implemented) | | Item identification | Gemini Vision API | | Product photo generation | Nano Banana API | | Browser automation | Browser Use (Playwright) | | eBay comp research | Browser Use β€” read only | | Depop listing creation | Browser Use β€” form population, pauses at submit | | Hosting | Render | --- ## **6\. Agent Architecture** Four distinct uAgents registered on Agentverse. Each has a single responsibility. All implement the Chat Protocol for Fetch.ai deliverable compliance. ### **Agent 1 β€” Vision Agent** **Responsibility:** Item identification \+ clean photo generation **Inputs:** Raw photo uploaded by user **Actions:** * Sends photo to Gemini Vision API * Extracts: item name, brand, model, condition, relevant keywords * Attaches confidence score to identification * Sends identified item \+ original photo to Nano Banana API * Nano Banana generates clean white-background product photo **Outputs:** Item identification object `{name, brand, model, condition, keywords, confidence}` \+ clean product photo URL **Triggers next:** Research Agent \+ Pricing Agent (in parallel) --- ### **Agent 2 β€” Research Agent** **Responsibility:** Real market comp data from eBay **Inputs:** Item identification object from Vision Agent **Actions:** * Launches Browser Use session (headed Chromium) * Navigates to eBay sold listings search * Searches `"{brand} {model}" sold listings` * Filters: last 90 days, condition match * Extracts: sale price, condition, date sold, listing title for top 10 results * Handles dynamic page load, infinite scroll if needed * Falls back to Mercari if eBay blocks session **Outputs:** Array of comp objects `[{price, condition, date, title}]` **Triggers next:** Pricing Agent (if not already running) --- ### **Agent 3 β€” Pricing Agent** **Responsibility:** Profit margin computation \+ eBay listing breakdown **Inputs:** Comp array from Research Agent **Actions:** * Filters outliers (\>2 standard deviations from mean) * Computes median sale price (not mean β€” more robust to outliers) * Applies Depop 10% fee, estimated shipping * Computes net profit assuming $0–15 thrift store purchase price range * Generates complete eBay listing breakdown: title, description, recommended price, category, condition notes **Outputs:** `{median_price, recommended_listing_price, estimated_profit, ebay_listing_breakdown}` **Triggers next:** Listing Agent --- ### **Agent 4 β€” Listing Agent** **Responsibility:** Depop form population via Browser Use **Inputs:** Item identification, clean photo URL, recommended listing price from Pricing Agent **Actions:** * Launches Browser Use session with pre-warmed logged-in Depop account * Navigates to listing creation flow * Populates sequentially (dynamic form β€” order matters): 1. Upload clean product photo (Nano Banana output) 2. Select category β†’ subcategory 3. Enter title (generated from brand \+ model \+ condition) 4. Enter description (generated from keywords \+ comp data) 5. Enter price (recommended\_listing\_price) 6. Select condition 7. Enter size if applicable * **Pauses at final submit button β€” does NOT post automatically** **Outputs:** Depop form fully populated, awaiting user confirmation to post --- ## **7\. Agent Sequencing** User uploads photo ↓ \[Vision Agent\] Gemini Vision β†’ item identification Nano Banana β†’ clean product photo ↓ β”Œβ”€β”€β”€β”€β”΄β”€β”€β”€β”€β” ↓ ↓ \[Research\] \[Pricing begins with Vision data\] eBay comps preliminary price estimate ↓ ↓ β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”˜ ↓ \[Pricing Agent\] Final median, profit margin, eBay breakdown ↓ \[Listing Agent\] Depop form populated ↓ Pause at submit β†’ user reviews β†’ clicks post Research and Pricing run in parallel after Vision completes. Listing fires only after Pricing completes. --- ## **8\. User Flow** 1. User opens the app 2. User uploads a photo of a thrift store item 3. App displays real-time agent activity feed β€” agents light up sequentially as they activate 4. Vision Agent: item identified, confidence score shown, clean product photo appears 5. Research Agent: eBay comps appear in real time as they're scraped 6. Pricing Agent: median price computed, profit margin displayed prominently 7. Listing Agent: Depop form shown populating field by field 8. Final state: two-panel results view β€” left shows agent summary \+ clean photo, right shows eBay breakdown card \+ Depop form preview with "Post to Depop" button 9. User clicks "Post to Depop" β€” Browser Use clicks submit --- ## **9\. UI Layout** ### **Single page, two panels** **Left Panel β€” Agent Activity Feed** * Four agent cards, each with status indicator (idle β†’ active β†’ complete) * Live log of what each agent is doing as it works * Item identification card: name, brand, model, condition, confidence score * Clean Nano Banana product photo displayed prominently **Right Panel β€” Results** * **Top section:** Profit summary card * Recommended listing price (largest text on screen) * Estimated profit range based on thrift store purchase price * Median eBay comp price * Number of comps analyzed * **Bottom left:** eBay breakdown card * Suggested title, description, price, category * Raw comp table: price, condition, date sold (top 5 shown) * "Copy to eBay" button * **Bottom right:** Depop listing preview * Shows all populated form fields * Clean product photo thumbnail * "Post to Depop" CTA button (prominent, primary action) ### **Design notes** * Dark theme * Agent activity feed uses animated status indicators β€” designed component, not a terminal log * Profit number is the largest text element on screen * Real-time updates via SSE β€” no polling, no page refresh * Mobile-responsive (stretch goal β€” Best Mobile Hack) --- ## **10\. Fetch.ai Compliance** **Required deliverables:** * ASI:One Chat session URL β€” agents implement Chat Protocol, demonstrable via ASI:One text interface * Agentverse profile URLs for all 4 agents β€” each registered independently with README and keywords * Public GitHub repo \+ demo video on Devpost **Architecture note:** ASI:One is not the user-facing entry point. The custom Next.js frontend is. Agents are registered on Agentverse with Chat Protocol implemented, satisfying the discoverability requirement. For Fetch.ai judging specifically, demonstrate text-based agent interaction via ASI:One as secondary proof alongside the custom frontend demo. --- ## **11\. Demo Script (3 minutes)** **0:00–0:20 β€” Setup** "We're standing in a thrift store. Found this item. Don't know if it's worth buying to flip. Let's find out in 60 seconds." Upload photo of pre-selected demo item (Air Jordan 1s or equivalent high-comp item). **0:20–0:50 β€” Vision Agent** Watch Vision Agent activate. Item identified: "Air Jordan 1 Retro High OG, Good condition." Confidence: 94%. Clean white-background product photo appears automatically. **0:50–1:30 β€” Research \+ Pricing** Research Agent opens eBay live. Comps appear in real time: "$145, $152, $138..." Pricing Agent computes: "Median $147. Recommended listing price $139. Estimated profit: $124–139 after fees." **1:30–2:30 β€” Listing Agent** Depop opens in Browser Use session visible to judges. Watch it populate live: photo uploads, category selected, title typed, description fills in, price entered. Form is complete. Pause at submit. **2:30–3:00 β€” Close** "Every field populated. One click to post." Click post. "That's the entire flipping workflow β€” identification, research, pricing, listing β€” fully automated. What used to take 30 minutes took 90 seconds." --- ## **12\. Pre-Hackathon Checklist (Tonight)** * Create Depop seller account, verify email, complete profile * Confirm Depop allows new accounts to create listings immediately (no new account restrictions) * Select demo item β€” Air Jordan 1s or equivalent * Verify demo item has abundant eBay sold comps * Verify demo item has meaningful profit margin ($50+) at thrift store prices * Obtain Nano Banana API credentials * Obtain Gemini API key * Set up Agentverse account, confirm uAgent deployment works * Spike on Browser Use β€” specifically test file upload and form population * Pre-warm Depop session in Browser Use (logged in, ready to go) * Register .Tech domain --- ## **13\. Risk Register** | Risk | Severity | Mitigation | | ----- | ----- | ----- | | eBay blocks Browser Use session during demo | High | Fall back to Mercari; pre-cache comps for demo item as absolute fallback | | Depop form structure changes or breaks | High | Map form manually pre-hackathon; hardcode field sequence for demo item category | | File upload via Browser Use fails on Depop | High | Test specifically tonight; have photo pre-staged at known absolute path | | Gemini misidentifies demo item | Medium | Use unambiguous item (Air Jordans); show confidence score as transparency feature | | Nano Banana photo quality poor | Medium | Pre-generate clean photo for demo item as fallback; display original if generation fails | | Agent sequencing race condition | Medium | Enforce strictly sequential pipeline; no agent starts without confirmed prior completion signal | | Depop new account listing restriction | High | Create account tonight, attempt test listing creation to confirm permissions | | Browser Use learning curve (no prior experience) | High | Person assigned to Browser Use spikes on it first β€” before writing any other code | | SSE connection drops during demo | Low | Implement reconnection logic; test on demo machine specifically | --- ## **14\. Scope Tiers** ### **Must ship β€” demo lives or dies on these** * Vision Agent: Gemini Vision identification \+ Nano Banana clean photo * Research Agent: eBay sold comp scraping via Browser Use * Pricing Agent: median price computation \+ profit margin \+ eBay breakdown * Listing Agent: Depop form population via Browser Use, pauses at submit * Frontend: two-panel layout with agent activity feed \+ results * SSE: real-time agent status updates pushing to frontend ### **Should ship β€” meaningfully improves submission quality** * Fetch.ai Agentverse registration \+ Chat Protocol for all 4 agents * eBay listing breakdown card with "Copy to eBay" button * Outlier filtering and condition matching in Pricing Agent * Mercari fallback in Research Agent * Polished animations on agent activity feed for Best UI/UX track ### **Stretch β€” only if must-ship and should-ship are complete** * React Native wrapper for Best Mobile Hack * Multi-item batch processing * User confirmation step between Vision Agent and downstream agents * Additional marketplace targets (Poshmark, Facebook Marketplace)